NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

The Key to Effective UDF Optimization: Before Inlining, First Perform Outlining

https://doi.org/10.14778/3696435.3696436

Arch, Samuel; Liu, Yuchen; Mowry, Todd C; Patel, Jignesh M; Pavlo, Andrew (September 2024, Proceedings of the VLDB Endowment)

Although user-defined functions (UDFs) are a popular way to augment SQL's declarative approach with procedural code, the mismatch between programming paradigms creates a fundamental optimization challenge. UDF inlining automatically removes all UDF calls by replacing them with equivalent SQL subqueries. Although inlining leaves queries entirely in SQL (resulting in large performance gains), we observe that inlining the entire UDF often leads to sub-optimal performance. A better approach is to analyze the UDF, deconstruct it into smaller pieces, and inline only the pieces that help query optimization. To achieve this, we propose UDF outlining, a technique to intentionally hide pieces of a UDF from the optimizer, resulting in simpler UDFs and significantly faster query plans. Our implementation (PRISM) demonstrates that UDF outlining improves performance over conventional inlining (on average 1.29× speedup for DuckDB and 298.73× for SQL Server) through a combination of more effective unnesting, improved data skipping, and by avoiding unnecessary joins.
more » « less
Full Text Available
Database Gyms

Lim, Wan Shen; Butrovich, Matthew; Zhang, William; Crotty, Andrew; Ma, Lin; Xu, Peijing; Gehrke, Johannes; Pavlo, Andrew (January 2023, Conference on Innovative Data Systems Research)

In the past decade, academia and industry have embraced machine learning (ML) for database management system (DBMS) automation. These efforts have focused on designing ML models that predict DBMS behavior to support picking actions (e.g., building indexes) that improve the system's performance. Recent developments in ML have created automated methods for finding good models. Such advances shift the bottleneck from DBMS model design to obtaining the training data necessary for building these models. But generating good training data is challenging and requires encoding subject matter expertise into DBMS instrumentation. Existing methods for training data collection are bespoke to individual DBMS components and do not account for (1) how workload trends affect the system and (2) the subtle interactions between internal system components. Consequently, the models created from this data do not support holistic tuning across subsystems and require frequent retraining to boost their accuracy. This paper presents the architecture of a database gym, an integrated environment that provides a unified API of pluggable components for obtaining high-quality training data. The goal of a database gym is to simplify ML model training and evaluation to accelerate autonomous DBMS research. But unlike gyms in other domains that rely on custom simulators, a database gym uses the DBMS itself to create simulation environments for ML training. Thus, we discuss and prescribe methods for overcoming challenges in DBMS simulation, which include demanding requirements for performance, simulation fidelity, and DBMS-generated hints for guiding training processes.
more » « less
Full Text Available
Litmus: Towards a Practical Database Management System with Verifiable ACID Properties and Transaction Correctness

https://doi.org/10.1145/3514221.3517851

Xia, Yu; Yu, Xiangyao; Butrovich, Matthew; Pavlo, Andrew; Devadas, Srinivas (June 2022, Proceedings of the 2022 International Conference on Management of Data)

Full Text Available
Tastes Great! Less Filling! High Performance and Accurate Training Data Collection for Self-Driving Database Management Systems

https://doi.org/10.1145/3514221.3517845

Butrovich, Matthew; Lim, Wan Shen; Ma, Lin; Rollinson, John; Zhang, William; Xia, Yu; Pavlo, Andrew (June 2022, Proceedings of the 2022 International Conference on Management of Data)

Full Text Available
Spitfire: A Three-Tier Buffer Manager for Volatile and Non-Volatile Memory

https://doi.org/10.1145/3448016.3452819

Zhou, Xinjing; Arulraj, Joy; Pavlo, Andrew; Cohen, David (January 2021, Proceedings of the ACM SIGMOD International Conference on Management of Data)

The design of the buffer manager in database management systems (DBMSs) is influenced by the performance characteristics of volatile memory (i.e., DRAM) and non-volatile storage (e.g., SSD). The key design assumptions have been that the data must be migrated to DRAM for the DBMS to operate on it and that storage is orders of magnitude slower than DRAM. But the arrival of new non-volatile memory (NVM) technologies that are nearly as fast as DRAM invalidates these previous assumptions.Researchers have recently designed Hymem, a novel buffer manager for a three-tier storage hierarchy comprising of DRAM, NVM, and SSD. Hymem supports cache-line-grained loading and an NVM-aware data migration policy. While these optimizations improve its throughput, Hymem suffers from two limitations. First, it is a single-threaded buffer manager. Second, it is evaluated on an NVM emulation platform. These limitations constrain the utility of the insights obtained using Hymem. In this paper, we present Spitfire, a multi-threaded, three-tier buffer manager that is evaluated on Optane Persistent Memory Modules, an NVM technology that is now being shipped by Intel. We introduce a general framework for reasoning about data migration in a multi-tier storage hierarchy. We illustrate the limitations of the optimizations used in Hymem on Optane and then discuss how Spitfire circumvents them. We demonstrate that the data migration policy has to be tailored based on the characteristics of the devices and the workload. Given this, we present a machine learning technique for automatically adapting the policy for an arbitrary workload and storage hierarchy. Our experiments show that Spitfire works well across different workloads and storage hierarchies.
more » « less
Full Text Available
Make your database system dream of electric sheep: towards self-driving operation

https://doi.org/10.14778/3476311.3476411

Pavlo, Andrew; Butrovich, Matthew; Ma, Lin; Menon, Prashanth; Lim, Wan Shen; Van Aken, Dana; Zhang, William (July 2021, Proceedings of the VLDB Endowment)

Full Text Available
Filter Representation in Vectorized Query Execution

https://doi.org/10.1145/3465998.3466009

Ngom, Amadou; Menon, Prashanth; Butrovich, Matthew; Ma, Lin; Lim, Wan Shen; Mowry, Todd C.; Pavlo, Andrew (June 2021, DAMON)

Full Text Available
Taurus: lightweight parallel logging for in-memory database management systems

https://doi.org/10.14778/3425879.3425889

Xia, Yu; Yu, Xiangyao; Pavlo, Andrew; Devadas, Srinivas (October 2020, Proceedings of the VLDB Endowment)
null (Ed.)
Existing single-stream logging schemes are unsuitable for in-memory database management systems (DBMSs) as the single log is often a performance bottleneck. To overcome this problem, we present Taurus, an efficient parallel logging scheme that uses multiple log streams, and is compatible with both data and command logging. Taurus tracks and encodes transaction dependencies using a vector of log sequence numbers (LSNs). These vectors ensure that the dependencies are fully captured in logging and correctly enforced in recovery. Our experimental evaluation with an in-memory DBMS shows that Taurus's parallel logging achieves up to 9.9X and 2.9X speedups over single-streamed data logging and command logging, respectively. It also enables the DBMS to recover up to 22.9X and 75.6X faster than these baselines for data and command logging, respectively. We also compare Taurus with two state-of-the-art parallel logging schemes and show that the DBMS achieves up to 2.8X better performance on NVMe drives and 9.2X on HDDs.
more » « less
Full Text Available
MB2: Decomposed Behavior Modeling for Self-Driving Database Management Systems

https://doi.org/10.1145/3448016.3457276

Ma, Lin; Zhang, William; Jiao, Jie; Wang, Wuwen; Butrovich, Matthew; Lim, Wan Shen; Menon, Prashanth; Pavlo, Andrew (June 2021, SIGMOD)

Full Text Available
Mainlining databases: supporting fast transactional workloads on universal columnar data file formats

https://doi.org/10.14778/3436905.3436913

Li, Tianyu; Butrovich, Matthew; Ngom, Amadou; Lim, Wan Shen; McKinney, Wes; Pavlo, Andrew (December 2020, Proceedings of the VLDB Endowment)

The proliferation of modern data processing tools has given rise to open-source columnar data formats. These formats help organizations avoid repeated conversion of data to a new format for each application. However, these formats are read-only, and organizations must use a heavy-weight transformation process to load data from on-line transactional processing (OLTP) systems. As a result, DBMSs often fail to take advantage of full network bandwidth when transferring data. We aim to reduce or even eliminate this overhead by developing a storage architecture for in-memory database management systems (DBMSs) that is aware of the eventual usage of its data and emits columnar storage blocks in a universal open-source format. We introduce relaxations to common analytical data formats to efficiently update records and rely on a lightweight transformation process to convert blocks to a read-optimized layout when they are cold. We also describe how to access data from third-party analytical tools with minimal serialization overhead. We implemented our storage engine based on the Apache Arrow format and integrated it into the NoisePage DBMS to evaluate our work. Our experiments show that our approach achieves comparable performance with dedicated OLTP DBMSs while enabling orders-of-magnitude faster data exports to external data science and machine learning tools than existing methods.
more » « less
Full Text Available

« Prev Next »

Search for: All records